Papers with jailbreak attacks

3 papers
Dagger Behind Smile: Fool LLMs with a Happy Ending Story (2025.findings-emnlp)

Copied to clipboard

Challenge: Large Language Models (LLMs) have attracted significant attention from jailbreak attacks . existing manual designs are either easily detectable or require intricate interactions with LLMs.
Approach: They propose a happy ending attack that wraps up a malicious request in a scenario template .
Outcome: The proposed attack wraps up a malicious request in a scenario template involving a positive prompt formed mainly via a happy ending, fooling LLMs into jailbreaking either immediately or at a follow-up malicious request.
Virtual Context Enhancing Jailbreak Attacks with Special Token Injection (2024.findings-emnlp)

Copied to clipboard

Challenge: Existing jailbreak attacks target the two phases of user interaction: prompt input and model computation.
Approach: They propose a new tool that leverages special tokens to improve jailbreak attacks . they found that the tool can increase success rates of existing jailbreak methods by 40% .
Outcome: The proposed solution can improve success rates of four widely used jailbreak methods by approximately 40% across various LLMs.
MTSA: Multi-turn Safety Alignment for LLMs through Multi-round Red-teaming (2025.acl-long)

Copied to clipboard

Challenge: Existing jailbreak techniques rely on single-round interactions, pro-Corresponding author.
Approach: They propose a multi-turn safety alignment framework to address the challenge of securing large language models in multi-round interactions.
Outcome: The proposed framework exhibits state-of-the-art attack capabilities while improving safety performance on safety benchmarks.

What is GenGO?

GenGO is an NLP powered publication search system. It currenctly indexes 30k+ papers from ACL Anthology, and implements multi-aspect summarization, semantic search, and more!

Information

About
Limitations